<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://sokwe.janegoodall.org/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=IanGilby</id>
	<title>sokwedb - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://sokwe.janegoodall.org/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=IanGilby"/>
	<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/wiki/Special:Contributions/IanGilby"/>
	<updated>2026-09-30T21:39:54Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.35.6</generator>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=815</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=815"/>
		<updated>2026-09-04T22:36:17Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=814</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=814"/>
		<updated>2026-09-03T22:23:31Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
&lt;br /&gt;
Will ask PIs about insect issues and Unrecorded/none/na&lt;br /&gt;
&lt;br /&gt;
&amp;quot;chipukizi&amp;quot; is correct. Do I change in all combinations, e.g. Chipukizi;majani?&lt;br /&gt;
same for utomvi&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=813</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=813"/>
		<updated>2026-09-03T21:32:24Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* SOLUTION */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=812</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=812"/>
		<updated>2026-08-25T23:07:10Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* SOLUTION */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=811</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=811"/>
		<updated>2026-08-24T23:18:35Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#52) There are follow_arrivals that are almost duplicates */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=786</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=786"/>
		<updated>2026-08-18T22:22:15Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=785</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=785"/>
		<updated>2026-08-18T22:14:57Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#99) There are GROOMINGS rows for which the extractedBy values is empty. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=784</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=784"/>
		<updated>2026-08-18T22:05:32Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=783</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=783"/>
		<updated>2026-08-18T21:37:30Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=782</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=782"/>
		<updated>2026-08-18T21:04:42Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=781</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=781"/>
		<updated>2026-08-18T20:53:06Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=780</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=780"/>
		<updated>2026-08-18T20:40:33Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#58) OTHER_SPECIES duplicate keys */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=778</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=778"/>
		<updated>2026-08-18T19:02:34Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#61) Invalid follow date/focals pairs in follow_arrival */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=777</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=777"/>
		<updated>2026-08-18T18:28:23Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#57) GROOM_BOUT duplicate keys */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=776</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=776"/>
		<updated>2026-08-18T17:41:36Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#53) There are follow_arrivals where the arriving chimp does not exist */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=775</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=775"/>
		<updated>2026-08-18T17:40:22Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* Solution */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=774</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=774"/>
		<updated>2026-08-14T20:11:04Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#49) There are follow_arrivals where females that are too old have a cycle code of U */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=773</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=773"/>
		<updated>2026-08-14T19:58:09Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#47) There are follow_arrivals where females that are too young have a cycle code of U */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=772</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=772"/>
		<updated>2026-08-14T19:47:24Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#105) There is a comm member log record that is missing a description. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=771</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=771"/>
		<updated>2026-08-14T19:46:57Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* Bad data */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=725</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=725"/>
		<updated>2026-07-07T21:11:28Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/schema&amp;gt;.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#98) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=724</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=724"/>
		<updated>2026-07-07T21:10:25Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* SOLUTION */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/schema&amp;gt;.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#98) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=722</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=722"/>
		<updated>2026-07-06T19:33:44Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/schema&amp;gt;.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#98) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=721</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=721"/>
		<updated>2026-07-06T18:40:42Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* SOLUTION */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/schema&amp;gt;.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#98) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=678</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=678"/>
		<updated>2026-07-03T18:49:19Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=677</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=677"/>
		<updated>2026-07-03T18:46:40Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=676</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=676"/>
		<updated>2026-07-03T18:46:22Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=675</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=675"/>
		<updated>2026-07-03T18:45:51Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=674</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=674"/>
		<updated>2026-07-03T18:44:41Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#92) There are food_bout rows that do not have a corresponding follow. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=673</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=673"/>
		<updated>2026-07-03T18:43:27Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#92) There are food_bout rows that do not have a corresponding follow. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=672</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=672"/>
		<updated>2026-07-03T18:42:19Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=671</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=671"/>
		<updated>2026-07-03T18:41:40Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=670</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=670"/>
		<updated>2026-07-03T18:40:36Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=669</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=669"/>
		<updated>2026-07-03T18:38:34Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#88) There are food_bout rows for which there is not a corresponding food name. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=668</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=668"/>
		<updated>2026-07-03T18:37:07Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ==&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=667</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=667"/>
		<updated>2026-07-03T18:34:43Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=666</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=666"/>
		<updated>2026-07-03T18:31:22Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#85) There are numerous instances of local food names that translate to multiple scientific food names. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=665</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=665"/>
		<updated>2026-07-03T18:29:48Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#84) Some follows have a community with trailing spaces */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=664</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=664"/>
		<updated>2026-07-03T18:28:58Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=663</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=663"/>
		<updated>2026-07-03T18:27:08Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=662</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=662"/>
		<updated>2026-07-03T18:20:13Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN OR VOLUNTEER TO FIX BY HAND&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=661</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=661"/>
		<updated>2026-07-03T18:19:03Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=660</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=660"/>
		<updated>2026-07-03T18:17:05Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=659</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=659"/>
		<updated>2026-07-03T18:15:00Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=658</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=658"/>
		<updated>2026-07-03T18:14:42Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=657</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=657"/>
		<updated>2026-07-03T18:14:03Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=656</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=656"/>
		<updated>2026-07-03T18:13:05Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAVE A LARGER DISCUSSION ABOUT HOW TO DEAL WITH &amp;#039;MALES&amp;#039;, &amp;#039;GROUP&amp;#039;, ETC&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=655</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=655"/>
		<updated>2026-07-03T18:10:38Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=654</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=654"/>
		<updated>2026-07-03T18:09:00Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=653</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=653"/>
		<updated>2026-07-03T18:06:17Z</updated>

		<summary type="html">&lt;p&gt;IanGilby: /* * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== * (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; 98 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; 84 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; 6 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== * (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
== * (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 363 AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>IanGilby</name></author>
	</entry>
</feed>